Back

Molecular Systems Biology

Springer Science and Business Media LLC

Preprints posted in the last 7 days, ranked by how well they match Molecular Systems Biology's content profile, based on 162 papers previously published here. The average preprint has a 0.12% match score for this journal, so anything above that is already an above-average fit.

1
Audited vibe coding suggests partial fetal-like convergence of tumor proteomes

Meyer, J. G.

2026-08-31 cancer biology 10.64898/2026.08.26.745609 medRxiv
Top 0.1%
12.6%
Show abstract

The balance between how much human tumors recapitulate fetal tissue programs versus lose adult tissue identity remains unresolved. I used audited vibe coding, a human-mediated, cross-model critique-and-refinement workflow, to re-analyze a public pan-cancer proteomic atlas. A primary large language model wrote and executed the analysis under scientific direction, while a separate model family audited the code, outputs and claims; findings were returned for correction across seven versioned releases. Among 229 tumor-adjacent pairs in seven organs, tumor-minus-adjacent proteomic change partially aligned with reverse fetal-to-adult maturation (organ-balanced cosine, 0.240; 95% interval, 0.138 to 0.335), with positive alignment in 189 of 229 patients (82.5%). The organ-balanced projection coefficient was 0.195 (95% interval, 0.069 to 0.244), indicating movement along only part of the developmental distance. Although reverse maturation overlapped adult-identity loss, a positive developmental component remained after identity loss entered first (0.203; 95% interval, 0.129 to 0.239). Suppression of adult-high proteins contributed to more positive alignment than reactivation of fetal-high proteins. The vibe coding audits identified substantive defects. A common-mask correction reduced the matched-organ advantage from 0.074 to 0.059; a missing-value correction barely changed aggregate geometry but replaced 5 of the top 40 liver contributors; and coupled resampling repaired uncertainty accounting without changing patient scores. As with any single report, the "vibe reanalysis" biological results are candidate discoveries pending independent replication. The workflow is a single feasibility case, not a reliability benchmark, and shows how conversationally generated analysis can be made more inspectable when model-written code is treated as untrusted, versioned and subject to separate-model critique and executable checks.

2
Critical Fragility Emerges from Chromosomal Instability in Cancer

Zambelli, F.; D'Addese, G.; Marti-Baena, Q.; Sardanyes, J.; Aguade-Gorgorio, G.; Sole, R.

2026-09-01 cancer biology 10.64898/2026.08.31.748208 medRxiv
Top 0.4%
4.9%
Show abstract

Genomic instability is a major driver of tumor evolution, promoting diversification and adaptation while simultaneously increasing the accumulation of deleterious alterations. How tumor populations balance these opposing effects remains poorly understood. Here, we introduce a computational framework that explicitly represents diploid genomes, functional gene classes, point mutations, and chromosome-segregation errors in spatially constrained and well-mixed tumor populations. We identify a viability boundary separating sustained tumor expansion from instability-induced population collapse. Within the viable regime, mutation and selection generate a stable distribution of genomic-instability classes that is accurately captured by an analytical replicator--mutator description. Near the viability boundary, tumor dynamics exhibit prolonged extinction transients and strong sensitivity to stochastic fluctuations, with important differences between solid and liquid architectures. Chromosomal alterations further modify growth by creating transient benefits through increased gene dosage and genetic redundancy, while ultimately increasing genomic fragility. Finally, simulated interventions show that eliminating low-instability subpopulations or increasing the global mutational burden can displace tumors beyond their viability boundary and trigger irreversible collapse. These results identify genome instability as both an evolutionary advantage and an intrinsic vulnerability, providing a quantitative framework for developing therapies that exploit the limits of tumor evolution.

3
RECON infers regions of interest from H&E images and reconstructs whole-slide molecular profiles at single-cell resolution

Yang, X.; Hao, N.; Zhao, R.; Angel, S.; Tan, Y.; Lian, C. G.; Zhou, L.; Olson, D.; Yu, K.-H.; Ruiz de Luzuriaga, A.; Wan, G.

2026-09-01 bioinformatics 10.64898/2026.08.25.747122 medRxiv
Top 0.4%
4.3%
Show abstract

Spatial omics technologies resolve molecular expression and spatial architecture at single-cell resolution, but profiling whole slides remains costly. In practice, only a few regions of interest (ROIs) are profiled, leaving the rest of the tissue unmeasured. S2-omics was the first framework to unify ROI selection with out-of-ROI prediction, but it operates on superpixels rather than individual cells and predicts discrete cell types rather than continuous molecular profiles. Superpixel-based representations do not explicitly preserve cell boundaries, while categorical cell-type labels cannot quantify molecular expression within cells. Here we present RECON, a two-stage framework that performs ROI inference and whole-slide molecular reconstruction at single-cell resolution, predicting both continuous molecular profiles and discrete cell-type labels. In the first stage, RECON extracts morphological and microenvironmental features from individual cells to identify a representative ROI for spatially resolved single-cell molecular profiling. In the second stage, RECON trains deep learning models on molecular measurements acquired within the selected ROI and reconstructs transcriptomic or proteomic profiles for all remaining cells on the slide. Benchmarked against pathologist annotations, RECONs ROI selection outperforms the superpixel-based S2-omics approaches (IoU: 0.75 versus 0.64). For transcriptomics, refining the modeling unit from superpixels to single cells improves per-gene Pearson correlation by 22%. For proteomics, RECON surpasses the current state-of-the-art method, ROSIE, across all 16 markers, with a median per-cell Pearson correlation of 0.91 versus 0.84. Moreover, RECON delineates tumour boundaries and regions with distinct immune-cell densities, and highlights candidate tertiary lymphoid structures. Together, these results demonstrate that RECON enables informative ROI selection and whole-slide molecular reconstruction at single-cell resolution for both spatial transcriptomics and spatial proteomics.

4
Rational Control of Basal CAR Expression Improves Discrimination in Inducible T Cell Circuits

Hoces, D.; Ng, J.; Perez, J.; Hernandez-Lopez, R. A.

2026-08-31 synthetic biology 10.64898/2026.08.28.747722 medRxiv
Top 0.5%
4.1%
Show abstract

SynNotch-CAR circuits improve T cell specificity by coupling antigen recognition to inducible CAR expression. However, basal CAR expression without receptor activation, termed here as leakiness, can reduce the separation between killing of intended target cells and sparing of antigen-positive off-target cells, limiting target-cell discrimination. Here, we systematically quantified basal CAR expression for several synNotch-CAR designs and developed a coupled ordinary differential equation model to show that discrimination depends on basal output, CAR potency, and effector-to-target ratio. We introduced C-terminal tags such as fluorescent proteins, degron domains, endocytosis signals, and endoplasmic reticulum retention motifs as a strategy to reduce CAR leakiness. We found that fluorescent proteins and degron-containing tags reduced basal CAR surface expression while preserving antigen-induced CAR expression, improving discrimination of antigen-density sensing and combinatorial circuits in vitro. In xenograft models, fluorescent protein-tagged CARs improved discrimination by reducing activity against off-target cells while retaining activity against high-antigen tumors. Degron-containing constructs reduced basal CAR expression in vitro but showed suboptimal performance in vivo, revealing a trade-off between basal CAR suppression and induced CAR persistence. Together, these findings demonstrate that basal output expression is a key parameter for inducible genetic circuit designs and establish layered transcriptional and post-translational regulation as a strategy to improve the fidelity of inducible T cell circuits.

5
Modeling Joint Reference Regions for Omics Biomarkers in UK Biobank Proteomics

Pusparum, M.; Thas, O.; Ertaylan, G.

2026-09-04 health informatics 10.64898/2026.09.01.26361504 medRxiv
Top 0.6%
3.3%
Show abstract

Conventional univariate reference intervals (UniRIs) are widely used to identify abnormal biomarker values, but they evaluate each biomarker independently and do not account for coordinated deviations between biomarkers. We developed and evaluated a joint reference region (JRR) framework for plasma proteomics data using the Olink proteomics dataset generated by the UK Biobank Pharma Proteomics Project, covering approximately 3,000 plasma proteins. JRRs were estimated for selected protein pairs in a healthy reference subset, while UniRIs were estimated separately for individual proteins using the nonparametric method. Both approaches were then evaluated in ICD-defined disease subsets. Biomarker discovery revealed sparse and heterogeneous disease--protein associations, with some proteins recurring across multiple phenotypes and others showing more disease-specific patterns. The added value of JRRs varied across diseases and protein pairs. Across evaluated protein pairs, 56.5\% showed higher sensitivity under the JRR framework than the UniRI of the first protein, and 47.3\% showed higher sensitivity than the UniRI of the second protein. At the disease level, the median proportion of protein pairs with improved JRR sensitivity was 0.57. JRRs were most informative when univariate detection was limited but a subset of diseased observations was flagged only by the joint region. These findings suggest that JRRs provide a complementary approach to UniRIs by capturing abnormal joint biomarker configurations in high-dimensional proteomics data.

6
MechanoMaST - a multimodal pipeline for spatially registering mechanical and transcriptomic tissue data

Decker, L.; Olisov, D.; Schleussner, N.; Wiethoff, H.; Schmidt, T.; Nienhueser, H.; Pausch, T. M.; Korbel, J. O.; Diz-Munoz, A.

2026-08-31 biophysics 10.64898/2026.08.29.747727 medRxiv
Top 0.9%
2.6%
Show abstract

Spatial-omics workflows enable molecular analysis within tissue spatial context. Despite the prognostic value of tissue stiffness, these approaches have not incorporated direct, mechanical measurements. This omission reflects several challenges, including sample requirements, low throughput, specialized equipment, and complex data registration. Here, we introduce mechanoMaST (mechanics mapped to spatial transcriptomics), the first workflow to combine absolute mechanical measurements with spatial-omics. It pairs atomic force microscopy-based nanoindentation stiffness maps with spatial transcriptomics maps from adjacent tissue cryosections. The two modalities are then computationally co-registered to enable direct spatial correlation at 100 um resolution, with mapping accuracy quantified through error propagation, providing ground-truth mechanical data directly linked to spatial gene expression. We demonstrate mechanoMaST in human colorectal cancer liver metastasis, generating a spatial resource from 10 patients and revealing a four-gene stiffness signature. mechanoMaST is readily adaptable to other tissues across development and disease, and extendable to additional spatial-omics modalities in adjacent sections.

7
Trans-Allosteric Activation Releases Distinct Conformational Traps in Kinase Heterodimers

Imamoto, A.; Wu, Y.; Shinobu, A.; Okada, M.

2026-09-01 biophysics 10.64898/2026.08.31.748385 medRxiv
Top 1.0%
2.4%
Show abstract

Protein kinases function as dynamic, mechanically coupled nodes, yet the conformational drivers of multimeric activation remain unclear. Here, we present AlloQuant, a computational suite that translates AlphaFold3 structural ensembles into quantitative metrics of kinase regulation, including internal network rigidity, metastable-state populations, and sub-angstrom conformational drivers. Applying AlloQuant to CDK1, we demonstrate that binding of the Cyclin B1 (CCNB1) cofactor mechanically decouples a hyper-rigid inactive kinase core, allowing activating phosphorylation (pT161) to subsequently re-impose localized tension on the catalytic machinery. Conversely, the C-terminal Src kinase (CSK) faces a distinct conformational trap. While nucleotide-free monomeric CSK spontaneously samples a pre-active geometry, ATP binding excludes the active C-In conformation in all but 1 of 225 models. We show that docking partner engagement overcomes this blockade. Autophosphorylation of SRC at the activation loop (Y419) redistributes SRC conformational states without altering bulk rigidity. This redistribution is structurally coupled to the conformational state of CSK via the regulatory spine, not the catalytic machinery. Rather than mechanically deforming CSK, SRC engagement acts by conformational selection, committing roughly a quarter of CSK molecules to a fully active state. Thus, trans-allosteric kinase activation operates by defining the accessible conformational landscape of the receiver kinase. That control is exerted through mechanical remodeling in cofactor-dependent complexes and through conformational selection in transient kinase-kinase heterodimers. These findings establish AlloQuant as a general framework for quantifying how a binding partner reshapes a kinase's conformational landscape, applicable across the kinome because it assigns landmarks by profile-HMM alignment.

8
MetaDome 2027: a comprehensively updated resource for aggregating missense variant evidence across homologous human protein domains

Wiel, L.; Ferraro, F.; Yu, J.; Zhen, J.; Nachun, D.; Mendez, R.; Reuter, C. M.; Cui, J. L.; Bonner, D. E.; Carter, J. N.; Marwaha, S.; van de Vorst, M.; Emami, S.; Kravets, E.; Neu, M. B.; van Ham, T. W.; Kleefstra, T.; Ashley, E. A.; Bernstein, J. A.; Montgomery, S. B.; Gilissen, C.; Wheeler, M. T.

2026-08-31 bioinformatics 10.64898/2026.08.26.747388 medRxiv
Top 1%
1.8%
Show abstract

The interpretation of missense variants remains a major challenge in clinical genetics. "Meta-domains" aggregate population and pathogenic variation across homologous Pfam domain instances in the human proteome, providing per-residue context for interpreting variants of uncertain significance (VUS). Our 2019 implementation, MetaDome, is widely used and named in clinical variant-classification guidelines. Here we present the MetaDome 2027 update, featuring a comprehensively updated dataset and GRCh38 support. The redesigned pipeline enables incremental updates of GENCODE, UniProtKB/Swiss-Prot, Pfam, gnomAD, and ClinVar while maintaining 100% sequence-identity gene-to-protein mapping. Annotated Pfam domain instances grew 14.9% from 71,419 to 82,069 and meta-domain-eligible Pfam families ([≥]2 human occurrences) by 73.3% from 3,334 to 5,778; Pfam domains are annotated to 92% of human proteins. Approximately 43% of mapped protein-coding nucleotides (14.3 million in GRCh38, 13.8 million in GRCh37) are in a meta-domain; in GRCh38 67.9% (37,692 of 55,548) of pathogenic or likely pathogenic ClinVar missense variants fall at such a position. We show how MetaDome helped reclassify a de novo missense VUS in RALA and identify 52,463 ClinVar missense VUS for which meta-domains supply otherwise unavailable pathogenic evidence. MetaDome is freely available at www.metadome.app.

9
JMod: Joint modeling of mass spectra for empowering multiplexed DIA proteomics

McDonnell, K.; Geiszler, D. J.; Wamsley, N.; Derks, J.; Sipe, S.; Cohen, Z. A.; Warinner, L. K.; Yeh, M.; Koo, E.; Leduc, A.; Zwang, T. J.; Specht, H.; Slavov, N.

2026-08-31 bioinformatics 10.1101/2025.05.22.655512 medRxiv
Top 1%
1.7%
Show abstract

Parallelization of data acquisition substantially increases the throughput of mass spectrometry-based proteomics. However, parallelization also increases the density of mass spectra and consequently the overlap between ions, frustrating their analysis. To improve sequence identification and quantification from such spectra, we developed an open-source software for Joint Modeling of mass spectra (JMod). JMod models overlapping peaks as linear superpositions of their components in both MS1 and MS2 space, which permits multiplexed DIA with smaller mass offsets to increase the multiplexing capacity and thus proteomics throughput for a given plexDIA tag. This enables 9-plexDIA using 2 Da offset PSMtags, increasing throughput 9-fold while preserving quantitative accuracy and coverage depth. Furthermore, we use JMod to deconvolve simultaneous labeling by mass tags and heavy amino acids, thus increasing the throughput of metabolic pulse experiments measuring protein synthesis and degradation rates in single cells from mouse liver. By supporting enhanced decoding of highly multiplexed DIA spectra, JMod provides an open and flexible software that increases the throughput of sensitive proteomics.

10
Ratiometric growth-rate control enables robust coexistence in competing microbial consortia

Barajas, C.

2026-08-31 synthetic biology 10.64898/2026.08.28.747825 medRxiv
Top 1%
1.7%
Show abstract

Maintaining a prescribed composition in engineered microbial consortia is difficult because small fitness differences can drive competitive exclusion. We study a two-strain consortium in continuous culture and develop a feedback architecture that regulates composition by selectively slowing the fast strain as a function of the population ratio. At the population level, we derive an idealized ratio-feedback law with a tunable positive coexistence equilibrium. We then propose a biomolecular realization using orthogonal quorum sensing, an sRNA-based ratiometric controller, and a ppGpp-mediated growth actuator. Exploiting the separation between slow population growth and faster intracellular controller dynamics, we use singular perturbation theory to show that, for sufficiently fast controller dynamics, the full implementation model inherits the coexistence equilibrium and its local stability properties from the reduced model. Numerical simulations validate the reduction and show how weaker timescale separation or loss of the assumed molecular regime degrades performance.

11
RegimeFormer: A Large Protein Model of Global Perturbation Regimes

Ma, S.; Chai, Y.; Wu, Y.; Zhang, Q.; Yuan, Y.; Zhao, K.; Chen, Z.; Wang, H.; Cao, S.; Yu, X.; Han, X.; Liu, Y.; Liu, Y.; Zhu, T.; Tao, D.

2026-08-30 bioinformatics 10.64898/2026.08.26.747182 medRxiv
Top 1%
1.7%
Show abstract

Protein language models organize sequence and structure at scale, but a global representation of how proteins respond to mutation remains lacking. We present RegimeFormer, a large protein perturbation model coupled to RegimeAtlas, constructed by harmonizing and indexing 202,556,313 non-redundant protein sequences across the tree of life. A diversity-preserving one-million-protein subset provides the high-resolution training and inference layer, with 995,995 proteins yielding residue-level summaries across 407,048,356 residues and substitution-specific predictions available on demand. Across experimental deep mutational scanning, molecular benchmarks, structural confidence and evolutionary constraint, RegimeFormer identifies reproducible protein-level perturbation regimes that organize residue fragility, adaptability and predictive uncertainty. Regime conditioning improves substitution-specific prediction, with the largest relative gains under unseen-protein, unseen-family and low-homology evaluation. RegimeFormer-derived molecular priors further improve downstream transcriptomic and drug-response modelling. Together, RegimeFormer and RegimeAtlas provide a scalable framework for mapping, predicting and querying protein perturbation landscapes across global sequence space.

12
Hormone oscillations preserve cellular responsiveness to future physiological demands

Greenwood, M.; Drube, J.; Hoffmann, C.; Li, P.

2026-08-31 systems biology 10.64898/2026.08.28.747949 medRxiv
Top 1%
1.7%
Show abstract

Living organisms must sense and adapt to physiological demands of varying intensity, requiring cells to remain responsive over time. While continuous changes in hormone concentrations communicate these demands, sustained stimulation desensitizes signaling, protecting cells from overstimulation but potentially blunting future responses. How cells preserve responsiveness remains unclear. Using epinephrine, a major mediator of stress responses, we show that natural ultradian oscillations provide a solution. Oscillatory, but not constant, hormone enabled receptor resensitization when hormone levels fell, preserving alertness to subsequent stress and tunability across intensities. Furthermore, oscillation supported coordinated responses among diverse cell types by more consistently maintaining responsiveness across hormone concentrations and receptor kinetics. Oscillations thus provide a general strategy by which endocrine systems retain protective desensitization while preserving responsiveness to future physiological demands.

13
Decoding Humoral Immunity During Acute MPXV Infection via Comprehensive Serological Analysis and Antigen-agnostic Monoclonal Antibody Profiling

Zhang, Y.; Fan, J.; Wang, J.; Jiang, N.; Wan, Y.; Meng, L.; Qi, W.; Cheng, X.; Luo, K.; Zhang, T.; Li, R.; Chen, H.; Zhao, R.; Ren, Y.; Zhang, W.; Zhu, Z.

2026-08-31 public and global health 10.64898/2026.08.21.26360138 medRxiv
Top 2%
1.7%
Show abstract

Dissecting the complexity of antibody responses in orthopoxvirus (OPXV) infected individuals is essential for elucidating protective mechanisms and identifying candidate protective immunogens. Here, we profiled the acute humoral response in 51 mpox cases, showing distinct IgG trajectories among multiple antigens alongside the rise of plasma neutralizing activities to plateau within 6 weeks after symptom onset. Utilizing a single-cell transcriptomic and BCR sequencing based antigen-agnostic mAb isolation workflow, we further generated monoclonal antibodies (mAbs) from 254 expanded peripheral B cell clones of 3 patients. We discerned 97 specific mAbs recognizing at least 12 different OPXV proteins via integrated screening approaches, which comprised neutralizing antibodies binding unconventional viral targets and antibodies exhibiting extraordinary in vitro and in vivo anti-OPXV effects. The number of OPXV-specific mAbs recovered per donor reflected the percentage of expanded clones among circulating B cells. More interestingly, we demonstrated that the inferred unmutated common ancestors (UCAs) of neutralizing antibody clones did not necessarily react with OPXV, implying that OPXV neutralizing antibodies might frequently originate from B cells previously activated by unknown antigens. Our work establishes an efficient workflow for antigen-agnostic isolation of pathogen specific mAbs and reveals previously unclarified features of antibody responses induced by acute MPXV infection.

14
Data coverage and model formulation reshape quantitative interpretations of bacterial transcriptional regulation

Kuo, S.-T. A.; Hsu, C.-P.; Chou, H.-H. D.

2026-09-01 systems biology 10.64898/2026.08.31.748186 medRxiv
Top 2%
1.4%
Show abstract

Thermodynamic models quantitatively describe interactions between transcription machinery and bacterial promoters. Contrary to conventional understanding, model analysis by Parisutham et al. (2025) attributes transcriptional inhibition by repressors to overstabilization of the RNA polymerase-promoter complex rather than prevention of its formation. Moreover, it suggests an inverse scaling relationship between basal promoter strength and transcriptional fold change, applicable to both repressor- and activator-mediated regulation. To reevaluate findings from this study, we systematically analyze empirical data and compare its framework with conventional thermodynamic models. In contrast to the inverse scaling relationship, data across multiple sources exhibit a peaked tradeoff between basal promoter strength and fold change, underscoring the importance of broad data coverage in revealing the full pattern required for reliable model inference. Furthermore, we identify the model assumption responsible for the apparent inverse scaling and misinterpretation of regulatory mechanisms. Relaxing this assumption enables the model to capture the peaked tradeoff and yield inferences consistent with established mechanisms of transcriptional repression and activation. We further derive a mathematical solution that connects basal expression to fold change for both repressor- and activator-regulated promoters. Our results underscore the importance of broad data coverage to avoid a blind-men-and-elephant interpretation and establish basal promoter strength as a key design parameter governing transcriptional regulation.

15
Substrate Profiling of RNF216 Uncovers a Translation-Linked OTUD4 Regulatory Axis

Wei, W.; Liu, R.; Zhang, J.; Liu, S.; Charles, A. J.; Asati, D. G.; Allen, Z. D.; Wright, D.; Peng, K.; Krekeler, E.; Mosammaparast, N.; Yin, J.; Mabb, A. M.

2026-08-30 neuroscience 10.64898/2026.08.26.747332 medRxiv
Top 2%
1.2%
Show abstract

Mutations in the E3 Ubiquitin (Ub) ligase RNF216 cause Gordon Holmes syndrome (GHS), a neurodegenerative disorder accompanied by neuroendocrine disruption. We developed an orthogonal ubiquitin transfer (OUT) platform to capture RNF216 substrates in neuronal cells and identified OTUD4, a deubiquitinating enzyme (DUB) mutated in GHS, and FMRP, a neuronal-enriched translational repressor. RNF216 predominantly synthesizes K6-linked Ub chains on OTUD4 to induce its degradation, forming donut-shaped structures in neurons. In return, OTUD4 removes the ubiquitination of RNF216 and FMRP. Analysis of RNF216 substrates revealed biological functions regulating protein synthesis, a shared function of the OTUD4-RNF216 substrate interaction network. Indeed, RNF216 expression increased protein synthesis rates in different cell types while Rnf216 deletion decreased dendritic development in neurons. Overall, our findings show that RNF216 and OTUD4 balance rates of protein synthesis and degradation and suggest GHS-related mutations in RNF216 or OTUD4 may offset this balance, triggering neurodegeneration.

16
A structural census links penultimate-residue class to N-terminal burial in human protein assemblies

Chang, Y.-H.

2026-09-01 biochemistry 10.64898/2026.08.31.748389 medRxiv
Top 2%
1.1%
Show abstract

Initiator-methionine excision is among the earliest protein modifications, yet its relationship to assembly geometry is unknown. Burial of the mature first residue was measured across 7,246 deposited human biological assemblies (22,291 chain-level observations; 1,191 proteins). Among 1,143 analyzable proteins, termini in MetAP-permissive penultimate-residue sequence classes were less often interface-engaged than termini in MetAP-nonpermissive classes (37.4% versus 47.4%; adjusted odds ratio 0.65, p = 7.2e-4). Curated processing annotations did not show a corresponding burial difference, and correlated residue properties preclude attributing the sequence-class association specifically to iMet removal. The analysis identified 264 interface-engaged MetAP-permissive candidates concentrated in cellular machines. In a fully recomputed conformer scan of deeply buried proteasome positions, modeled methionine accommodation was less favorable than at observed-methionine controls (median overlap -0.30 versus -1.12 angstrom, p = 0.0049), although most scoreable sites permitted a nonoverlapping placement. The census therefore reveals a graded structural constraint - not universal steric failure - and prioritizes complexes in which altered packing, assembly kinetics, lipidation or N-terminal methylation can be tested.

17
A thermodynamic framework for mapping elastic recoil mechanism across the human proteome

Desai, R.; Pople, D.; Musale, A.; Jain, S.; Sajjad, I.; Wittebort, R. J.; Koder, R. L.; Nanda, V.

2026-08-30 biophysics 10.64898/2026.08.28.747957 medRxiv
Top 2%
1.1%
Show abstract

The folding thermodynamics of proteins are dominated by two opposing forces, the loss in backbone entropy and the packing of hydrophobic groups. The same forces are major contributors to the extension thermodynamics of elastic proteins with the distinction that both processes act in concert, favoring the higher chain and solvent entropy of a relaxed conformation. The relative entropic contributions specify the recoil mechanism; human elastin recoil is primarily driven by hydrophobic forces, whereas fly resilin has a rubber-like mechanism driven by backbone entropy. Despite the importance of elastic proteins to tissue biomechanics, few have been identified, let alone characterized to the same extent as elastin and resilin. We develop a thermodynamic framework that maps proteins by sequence-derived estimates of extension-induced backbone and solvent entropy changes. Putative elastic proteins are proposed and classified by recoil mechanism based on estimated thermodynamic features. Proteins that map to elastic regions are overrepresented by the skin proteome. The set of predicted elastic domains is further extended by incorporating sequence context embedded in protein language models. Protein domains with distinct thermodynamic recoil mechanisms cluster on the latent space manifold. Some of these domains are anticipated to have roles within molecular machines, expanding the scope of elastic protein function beyond mechanical materials like elastin and resilin.

18
A novel framework leveraging non-causal associations reveals shared pathways linking inflammation and cancer risk

Yarmolinsky, J.; Cavallo, F. R.; Koskeridis, F.; Yu, X.; Bouras, E.; Richenberg, G.; Costantini, I.; Ray, D.; Woolf, B.; Karhunen, V.; Ellis, L.; Haycock, P. C.; Hemani, G.; Davey Smith, G.; Tsilidis, K. K.; Zuber, V.; McKay, J. D.; Dehghan, A.; Tzoulaki, I.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.30.26361622 medRxiv
Top 2%
1.1%
Show abstract

Confounding is a central challenge in observational studies. Here, we propose a framework for identifying confounders of two non-causally related traits by employing cross-trait pleiotropy analysis to detect genetic loci that affect both traits and multi-trait colocalisation to identify molecular phenotypes mediating these effects. We apply this approach to the analysis of C-reactive protein (CRP) - a non-specific marker of inflammation - and 10 inflammation-related cancers. In UK Biobank, higher pre-diagnostic CRP levels are associated with increased risk of multiple cancers, but bidirectional Mendelian randomization provides little evidence for a causal relationship. Cross-trait genetic analyses identify 92 loci with shared CRP-cancer effects including those with established roles in cancer and 50 novel loci such as RSPO3 (breast cancer) and GCKR (colorectal cancer). Integration with proteomic and single-cell transcriptomic data identified putative molecular mediators at 24 loci including plasma TLR1 levels in breast cancer and CD4+ T cell IRF5 expression in kidney cancer. Notably, 15 candidate effector genes encode targets of approved or investigational medications, including IL6, PDE4D, and CASP8, indicating potential opportunities for their repurposing for cancer prevention. The proposed approach provides a generalisable framework for leveraging non-causal phenotypic relationships to yield insights into disease mechanisms and therapeutic targets for disease prevention.

19
Harnessing Escherichia coli motility to engineer bacterial Voronoi patterns

Park, J. H.; Boni, E.; Hollo, G.; Schaerli, Y.

2026-09-01 synthetic biology 10.64898/2026.08.31.748246 medRxiv
Top 2%
1.1%
Show abstract

Cell motility drives spatial pattern formation across diverse biological systems. Here, we engineer Escherichia coli motility in semi-solid agar to control Voronoi patterns in two and three dimensions, partitioning space into regions closest to their respective inoculation seeds. Consistent with our reaction-diffusion model, we observed that collisions between expansion fronts generate either biomass depletion (''gaps'') or accumulation (''anti-gaps''), governed by the relative diffusion rates of bacteria and nutrients. By engineering strains with distinct expansion rates and tuneable motility, and by integrating these experimental data into a dynamic Voronoi model, we achieved precise control over pattern geometry. This enabled the generation of gaps with varying widths, curved boundaries, asymmetric structures, seedless regions, and complex composite patterns. Together, these findings establish bacterial Voronoi patterns as a programmable platform for engineering multicellular spatial organization, with potential applications in synthetic biology and materials science.

20
Chemi-Proteome Language Attention Network Empowers Fragment-Based Ligand Interactome and Binding Sites Discovery with Evidence

Liao, B.; He, J.; zhao, M.; Cui, X.; Cui, Y.; Dong, C.; Sun, H.; Zhang, L.; Zhang, J.

2026-08-30 bioinformatics 10.64898/2026.08.26.747036 medRxiv
Top 2%
1.1%
Show abstract

Deep learning has accelerated drug discovery, yet most existing models are trained using in vitro affinity datasets and consequently remain disconnected from the cellular context in which functional ligand-protein interactions occur. This limitation hinders the ability to reflect the complexity of native interactomes and characterize biological responses to molecular perturbation. Here we introduce C-PLANK (Chemi-Proteome Language Attention NetworK), a deep learning framework trained on fragment-protein interactions profiled directly in living cells using fully functionalized fragment (FFF) chemoproteomics. C-PLANK combines physicochemical embeddings with a bilinear attention network (BAN) to model both global cellular context and local residue-atom interactions, generating interpretable interaction fingerprints. Particularly, C-PLANK incorporates Cellular Interaction State Index (CISI), a systems-level evidential metric that contextualizes the biological plausibility of each predicted interaction against the global cellular interaction landscape. Across 431 ligand interactomes curated from eight independent chemoproteomic studies, C-PLANK consistently outperformed current state-of-the-art interaction prediction frameworks under both random and cold-protein evaluation settings. The inferred interaction fingerprints aligned with orthogonal evidence from structure-based pocket predictions, co-crystal structures, and cellular binding-site annotations. C-PLANK further generalized to unseen ligands. In a cellular target-focused discovery campaign, C-PLANK identified a previously unrecognized ligand that was subsequently advanced into an active chemical probe acting as a SIRT3 agonist in cellular assays. By learning directly from cellular chemoproteomics, C-PLANK moves beyond isolated interaction prediction toward cellular interaction-state modelling, establishing a computational foundation for future digital-twin frameworks in drug discovery.